Skip to content

test: assert Level-1 candidate membership and account for every rejection - #728

Open
zzylol wants to merge 1 commit into
issue-752from
test/promql-exact-function-coverage
Open

zzylol wants to merge 1 commit into
issue-752from
test/promql-exact-function-coverage

Conversation

@zzylol

@zzylol zzylol commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Review candidate correctness independently of cost ranking.

Before this PR: workload tests selected a synthetic-cost winner and mixed structural checks with ranking assertions. Valid Planner computation could also be rejected when a Backend binding lost its semantics.

After this PR: Level 1 enumerates the supported candidate inventory before pricing for ten individual queries and three ensembles: shared-rate, shared-quantiles, and the full workload. It checks typed DAG dependencies, windows, grouping, Rate sort expressions, materialization boundaries, dataset-bound SDS definitions, and shared producer identity. A complete Rate frontier may itself be a physical output; its identity program is checked explicitly.

There is no binding-defect allowlist. Unexpected compilation failures fail the suite. The strict fixture explicitly rejects heap candidates without certified accuracy evidence; the companion fixture proves those families bind when the required evidence is supplied. Every admission record remains unpriced.

Validation: all four Level 1 tests pass, including the wrong-sort-key mutation, certified heap admission, and ensemble checks. Admission reports and review instructions are updated. Complete candidate JSON/DOT and ensemble exports are generated as CI artifacts, not committed bundles. These exports support human review and do not constitute human approval.

Review sequence: #728 checks structure; #786 adds ERP/analytical costing; #742 checks synthetic-cost ranking; #775 installs selected candidates and executes the data plane. Candidate search has an explicit bounded scope. Real workload/resource measurements, online ERP feedback, and runtime replanning remain deferred.

@zzylol zzylol changed the title feat(promql): add executable exact bindings and full function smoke coverage test: verify issue 754 level-1 physical plans Sep 22, 2026

@zzylol zzylol left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Plesae actually run the asapplanner and asapquery-backend to get possible physical plan output, and I will review whether they are corerct.

readout: "sum",
root_operation: None,
},
"spatial-topk" => ExpectedPlan {

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should be some CMS/CS sketches with heap?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I ran the pinned ASAPPlanner/backend compiler and exported every compilable candidate. For topk by (label_0) (3, data), it emits an exact fallback and a backend-local CurrentSeries(data, by label_0) -> TopK(k=3) readout. The test now selects and asserts the local plan. CMS/CS heaps estimate frequency-heavy hitters, while this PromQL TopK ranks the current sample values and must retain the original series labels; using a frequency sketch directly would change the query. A certified candidate-membership sketch plus exact reranking could be a separate optimization.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed, and the enumerator already does this — the test was the thing that was silent.

Planner exposes a CountSketchWithHeap candidate for spatial-topk over the same current-series snapshot. It is enumerated, then rejected before pricing with selected summary readout has no certified accuracy guarantee, because this fixture supplies no distinct-item bound and no score separation. The exact Sort → Limit program is what binds.

family: None used to mean the test asserted nothing here. It now declares the shape explicitly:

"spatial-topk" => &[(
    "CountSketchWithHeap",
    Resolution::MustReject(PolicyReason::NoCertifiedGuarantee),
)],

So the heap candidate is required to be present and required to be refused for that specific reason — if Planner stops exposing it, or refuses it for a different reason, the test fails. admission/spatial-topk.admission.json records it.

A certified companion fixture is not derivable from this generator; see the reply on the quantile thread and the note in required_summary_shapes.

Comment thread control_plane/tests/issue754_level1.rs
Comment thread control_plane/tests/issue754_level1.rs
Comment thread control_plane/tests/issue754_level1.rs Outdated
root_operation: None,
},
"temporal-rate" => ExpectedPlan {
family: Some("Increase"),

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what's the difference between rate and increase? Should they be two nodes?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

They are separate physical roles here: the maintained Increase producer stores reset-aware counter state; the query DAG has an ExactReadout::Rate over that producer, applying the rate/window readout. The test now asserts that two-node producer/readout relationship rather than treating Increase as a Rate answer.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ExactKind has them as distinct variants (ASAPPlanner crates/types/src/post_asap/sketch.rs):

pub enum ExactKind { Sum, Count, Min, Max, Increase, Rate, IRate }

Increase is the counter-reset-aware delta; Rate is that delta divided by the window duration. So they are separate operators, not one node with a flag.

They are not two nodes in a plan, though. Every rate candidate in this fixture materializes as a single node — {"family":{"ExactAggregate":["Rate","Rate"]},"kind":"summary_agg","reduction":"PerEntity"} with a Readout{statistic:"Rate", logical_lookback_ms:"60000"} at the physical layer. The reset-aware delta is internal to the Rate accumulator rather than a separate Increase node feeding it.

Worth flagging: Increase appears in zero candidates across all ten queries, and so do Count, Min, Max and IRate. The ten-query workload only exercises Sum, Rate and quantiles. That is a real coverage gap in a branch named promql-exact-function-coverage — I think it belongs in a follow-up rather than here, but it should be a deliberate decision.

readout: "rate",
root_operation: None,
},
"grouped-rate" => ExpectedPlan {

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this should be, the inner node is rate, outer node is grouped sum exact aggregation.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed. The actual selected chain is Logical Aggregate(sum by label_0) -> ExactReadout(rate) -> ReadMaterialization(Increase). The test now asserts the inner rate readout, the outer exact grouped Sum, and their edge.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both placements are enumerated and both bind — the inventory is not committed to one.

  • grouped-rate-4: query-time grouped Sum over complete per-series Rate readouts. SummaryBuild(Sum) appears only in the query DAG.
  • grouped-rate-3: fixed-window precompute over complete per-series counter states. SummaryBuild{ExactAggregate:["Sum","Sum"]} runs at maintenance time; the query DAG is just Readout{statistic:"Sum"}.

In both, each series is finalized with PromQL Rate semantics before Sum groups the values — assert_native_grouped_rate checks that the Sum consumes a Readout{statistic:"Rate", logical_lookback_ms:"60000"}, so summing raw counters before per-series rates would fail.

The test now requires both placements to appear (grouped_rate_placements must contain {false, true}), rather than asserting one expected shape. Which of the two wins is #742;'s call, not this layer's.

One defect is visible here: a third precompute-Sum variant fails with Planner logical fragment does not match any original query subtree, and a fourth with a schema incompatibility. main is green on that path, so both are regressions from inside this stack. They are pinned in KNOWN_BINDING_DEFECTS with exact counts so they cannot be fixed or worsened silently.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Correction to my earlier reply on this thread.

I said the two bind_failed candidates here were regressions introduced inside the #737 → #728 stack, because main is green on these paths. That was wrong, and I should have diffed before asserting it.

residual_nodes is substantively identical on main. The only things this stack changed in that function are how the accuracy target is derived (previously hardcoded Exact) and the wording of the error message. main is green because its tests never plan these queries, not because the code is correct. The defect is older and wider than I claimed.

Root cause, now diagnosed rather than guessed. residual_nodes locates the query-time fragment by re-parsing every subtree of the original and comparing typed IR with ==. For this query the two sides differ in exactly one field:

residual   leaf scan columns: [ts, value, label_0]
candidate  leaf scan columns: [ts, value]

Everything else is identical: reduction, measures, output_names, range, source, predicates, time_index. Planner documents a PromQL leaf schema as usage-derived, "the (ts, value) floor + the labels the query references", and marks it closed: false because it deliberately does not enumerate the row. rate(data[1m]) carries label_0 only because the enclosing sum by (label_0) references it; re-parsing that subtree on its own cannot know that.

Fixed on main in #781 by comparing open leaf schemas by containment (a prefix, so positional ColumnIds keep their meaning) while everything else stays exact. The pre-existing false-match guards, which discriminate on matchers, range and metric name, all pass unchanged.

KNOWN_BINDING_DEFECTS and the review README have been corrected to say this.

readout: "rate",
root_operation: Some("aggregate"),
},
"grouped-temporal-sum" => ExpectedPlan {

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this should be fused by ASAPPlanner into one sum, over time and spatial dimensions.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The actual Planner candidate is currently Logical Aggregate(sum by label_0) -> ExactReadout(sum) -> per-entity Sum producer; it has not fused time and spatial aggregation into one grouped temporal Sum producer. I changed the expected contract to the fused producer, so the level-1 test now fails and names this gap. This comment remains an implementation blocker rather than being marked resolved.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The fused candidate exists and is the one that binds cheapest — grouped-temporal-sum-0 is a single operator over both dimensions:

{"kind":"summary_agg","family":{"ExactAggregate":["Sum","Sum"]},"reduction":{"Reduce":[2]}}

with "window_size":60 (the 1m temporal range) and "grouping_labels":{"labels":["label_0"]} (the spatial grouping) on that same node, "partitioning":"grouped". No separate per-entity stage.

The unfused shape is also enumerated (grouped-temporal-sum-3): summary_agg{Sum, reduction:"PerEntity"} → FinalizeExactAccumulator → a distinct Aggregate{sum, Reduce} node. Both are legal frontiers, so Level 1 keeps both and asserts neither as the winner.

On your earlier point about the 1-minute read: the grouped branch now checks readout_lookback_ms = 60_000 and that window_ms matches the materialization's stored_window_ms(), so a 5-second pane cannot pass as the 1-minute answer.

Also here: one candidate fails with Planner logical fragment does not match any original query subtree, recorded in KNOWN_BINDING_DEFECTS.

readout: "sum",
root_operation: Some("aggregate"),
},
"topk-rate" => ExpectedPlan {

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There should be inner node of rate, outter node for some sketches for topk

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The actual local chain is TopKSelection(k=3, by label_0) -> ExactReadout(rate) -> ReadMaterialization(Increase). It computes rate first and then selects TopK locally. A CMS/CS frequency heap is not itself an exact TopK of rate values; a sketch candidate stage would require a completeness certificate and exact rerank. The test asserts the current semantically correct local chain, and all candidate plans are exported for review.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That shape is enumerated — four candidates of it, in fact:

inner Readout{statistic:"Rate"} → outer KeyedSummaryBuild{Sketch:[CmsWithHeap | CountSketchWithHeap]} → KeyedReadout → Sort → Limit.

Two are query-time, two are fixed-window candidates that finalize Rate and build the heap during precompute. All four are rejected before pricing with no certified accuracy guarantee, so exact Sort → Limit binds.

The test now requires both families and that specific resolution:

"topk-rate" => &[
    ("CmsWithHeap", Resolution::MustReject(PolicyReason::NoCertifiedGuarantee)),
    ("CountSketchWithHeap", Resolution::MustReject(PolicyReason::NoCertifiedGuarantee)),
],

plus the pre-existing structural checks that the maintenance DAG contains Rate + KeyedSummaryBuild while the query DAG contains KeyedReadout and not KeyedSummaryBuild, and that ranking sorts on the finalized rate value rather than a timestamp or label.

readout: "rate",
root_operation: Some("top_k_selection"),
},
"quantile-ratio" => ExpectedPlan {

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this should be DDSketch / DDSketch

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed on the intended composition: q=0.9 and q=0.5 quantile-sketch readouts over the same per-series window, followed by local division. I added that target DAG contract. The current Planner emits only ExactFallback candidates, so the level-1 test remains red for this query; I am not treating fallback as a correct answer.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is what happens, and the test now says so instead of asserting nothing.

quantile-ratio-0 is the bound candidate: one summary_agg{Sketch:[DDSketch,...]} feeding two summary_estimate{Quantile} nodes (q=0.9 and q=0.5) into a binary/Div. So DDSketch / DDSketch, sharing a single producer.

KLL is attempted and discarded before admission — its composed guarantee does not satisfy EpsilonDelta{epsilon:0.01, delta:0.01} for either quantile. The other two candidates degrade to exact_fallback as a result. Exact (non-fallback) never appears.

The existing assertions already check the operand structure, the shared per-entity materialization at readout_lookback_ms = 60_000, and that both readouts come from a quantile-sketch family. What changed is that a KLL rejection for any reason other than an accuracy-target miss now fails the test rather than passing quietly.

@zzylol

zzylol commented Sep 22, 2026

Copy link
Copy Markdown
Contributor Author

I ran the pinned ASAPPlanner through the backend's real candidate enumeration and physical compiler for all ten issue #754 queries. The test exports both selected and every compilable candidate as JSON/DOT through the PR #736 renderer; the current CI run is https://github.com/ProjectASAP/ASAPQuery-backend/actions/runs/35736247732 (artifact issue754-level1-plans is uploaded even when the strict test fails).

Query Current cost-selected query DAG
spatial sum grouped Sum producer → Sum readout
spatial TopK backend-local CurrentSeries(data, by label_0) → TopK(k=3) readout
spatial quantile grouped DDSketch → q=0.9 readout
temporal sum per-series Sum → Sum readout
temporal quantile per-series DDSketch → q=0.9 readout
temporal rate per-series Increase → Rate readout
grouped rate per-series Increase → Rate readout → exact Sum by label_0
grouped temporal sum per-series Sum → Sum readout → exact Sum by label_0 (not the requested fused producer)
TopK rate per-series Increase → Rate readout → local TopK(k=3, by label_0)
quantile ratio ExactFallback (no local candidate)

The test now checks the specific DAG contracts and stays red for the two highlighted gaps. These are actual compiler outputs for the single-query level-1 snapshots with deterministic test costs, not measurements of production cost selection or data-plane execution. #742 is the separate runtime/semantic gate.

@zzylol zzylol changed the title test: verify issue 754 level-1 physical plans fix: preserve Planner families in issue 754 level-1 plans Sep 22, 2026
@zzylol zzylol changed the title fix: preserve Planner families in issue 754 level-1 plans refactor: align physical DAG, SDS, catalog, and runtime with Planner families Sep 22, 2026
@zzylol zzylol changed the title refactor: align physical DAG, SDS, catalog, and runtime with Planner families test: assert issue 754 level-1 physical DAGs for every query Sep 22, 2026
@zzylol
zzylol changed the base branch from main to fix/precompute-post-asap-dag September 22, 2026 16:16
@zzylol
zzylol changed the base branch from fix/precompute-post-asap-dag to refactor/backend-plan-split September 22, 2026 20:45
@zzylol
zzylol changed the base branch from refactor/backend-plan-split to refactor/typed-summary-plan-support September 24, 2026 12:56
@zzylol

zzylol commented Sep 26, 2026

Copy link
Copy Markdown
Contributor Author

For grouped-rate (sum by (label_0) (rate(data[1m]))), this Level-1 test fixes one particular split: a per-series Increase/Rate materialization, then query-time ExactReadout(rate) → Aggregate(sum by label_0). That is semantically valid, but the grouped Sum need not always remain on the query side. A precompute DAG could evaluate each series' reset-aware rate for a defined window/evaluation time, sum those rates by label_0, and materialize the grouped result. Summing raw counters before computing per-series rates would not be equivalent.

The planning objective is minimum execution cost for the workload, accounting for maintenance CPU, retained state, query CPU/read volume, update cadence, and sharing across queries. Hard-coding this split as the definition of plan correctness could reject a cheaper, semantically valid candidate. The current deterministic test costs also do not establish a production cost optimum.

Could we make the correctness contract assert rate-before-grouped-sum semantics, the window/evaluation-time and population contracts, and writer/read bindings independent of where the Sum executes? Then separately test that explicit workload/cost evidence selects the cheapest admitted executable candidate, and use Level 2 to compare the selected plan's results. If this shape is intentionally pinned only for this fixture, please label it as a fixture-specific selection expectation rather than a general correctness requirement.

@zzylol

zzylol commented Sep 26, 2026

Copy link
Copy Markdown
Contributor Author

Next query: grouped-temporal-sum (sum by (label_0) (sum_over_time(data[1m]))). I see a concrete Level-1 assertion gap here. The expected plan requires a grouped Sum materialization with window_size = 5 and a reduce read, but the spatial/grouped branch of assert_selected_plan does not check readout_lookback_ms = 60_000 (the assertion is only in the per-entity branch). A 5-second grouped pane is not itself the 1-minute query answer: the read must cover the complete 1-minute range, merge/reduce the right panes, and respect the evaluation-time boundary. Could we assert that temporal read/coverage contract explicitly for this case?

Separately, this test deliberately fails the current per-series Sum → readout → query-time grouped Sum plan and requires a fused grouped temporal Sum producer. Fusion is a valid candidate, but the per-series plan can also be semantically correct. Given the goal of minimum workload execution cost, I would test both as admissible when their coverage and bindings are valid, then use explicit workload/cost evidence to select the cheaper one. Please keep a fixed fused shape only if this is a fixture-specific cost-selection expectation, not the general correctness definition.

@zzylol

zzylol commented Sep 26, 2026

Copy link
Copy Markdown
Contributor Author

Concrete change proposed for this Level-1 acceptance test and the planning path (following the grouped-rate/grouped-temporal-sum comments):

  1. In ASAPPlanner candidate generation, enumerate legal materialization frontiers for the Post-ASAP DAG, including both (a) per-series window state → query-side grouped Sum and (b) per-series window computation → precompute grouped Sum → stored grouped output. Preserve the rate-before-sum ordering, evaluation-time/window semantics, labels, and maintenance lifecycle. A frontier is a candidate, not an instruction for the deployment compiler to move operators after selection.
  2. For each candidate, use Planner physical compilation to lower and validate both DAGs. Ask the backend for deployment feasibility (source/state bindings, pane coverage, historical retention, supported operators). Reject unsupported/incomplete candidates before pricing. The backend's DeploymentPlanCompiler then binds and publishes the selected physical DAGs; it must not silently substitute a different computation.
  3. Compare feasible candidates at workload scope, not one query at a time. Price update/maintenance work, stored bytes and retention, query execution/read cost weighted by recurrence, and producer sharing across queries. Use feat: install, execute and recover native candidates over dataset-bound SDS #761's scoped evidence/resource costs where applicable. Select the lowest-cost candidate satisfying accuracy and resource constraints; keep the chosen split and cost breakdown in the plan artifact for review.
  4. Split test: assert Level-1 candidate membership and account for every rejection #728's checks: a semantic/structural gate for every admitted candidate (including 1-minute coverage for grouped-temporal-sum), and a selection gate with at least two controlled cost fixtures that reverse which split wins. Assert that the selected DAG changes accordingly. Run test: rank individual and ensemble candidates with synthetic workload costs #742's local-result comparison for the selected alternatives. The current 1/2/1e12 synthetic pricing and fixed expected DAG can test a particular fixture, but cannot establish workload-optimal placement.

This seems consistent with Planner #462 owning deployment-independent physical lowering, #761 owning backend workload costing/evidence, and #737/#749 binding selected producer/read plans through stored output identity. Does this match the intended ownership?

@zzylol
zzylol force-pushed the test/promql-exact-function-coverage branch from 3159bdf to ef58ad2 Compare September 26, 2026 04:42
@zzylol
zzylol changed the base branch from refactor/typed-summary-plan-support to issue-752 September 26, 2026 04:43
@zzylol zzylol changed the title test: assert issue 754 level-1 physical DAGs for every query test: validate all supported physical candidate structures before pricing Sep 28, 2026
@zzylol zzylol changed the title test: validate all supported physical candidate structures before pricing test: validate individual and ensemble physical candidate DAGs before pricing Sep 28, 2026
zzylol added a commit that referenced this pull request Sep 28, 2026
…tion

Level 1 declared one expected plan per query while its stated purpose was to
enumerate the candidate inventory before pricing. A single expected shape is a
selection assertion in structural clothing, and four of the ten queries carried
no expectation at all.

Replace it with a membership contract. Each query declares the shapes the
inventory must expose and how this fixture resolves them: MustBind, or
MustReject with a named policy reason. The heap families Planner exposes for
spatial-topk and topk-rate are now required to be present *and* refused, which
is a positive assertion about the inventory rather than silence about it.

Separate policy refusals from defects. A rejection that is neither a declared
policy reason nor a recorded defect now fails the test. Three candidates fail
with "Planner logical fragment does not match any original query subtree" and
one with a schema incompatibility; main is green on these paths, so each is a
regression introduced inside the #737 -> #728 stack. They are listed in
KNOWN_BINDING_DEFECTS with exact occurrence counts, so a new instance fails and
a fixed one forces its entry to be deleted. All must be gone before #775
installs and executes these candidates.

Stop committing derived plans. candidates/ and ensembles/ were ~94k lines, are
fully reproducible from the test, and every identity in them changes when the
Planner pin moves; CI already uploads them on every run. Only the per-query
admission reports stay in the tree, and they now reproduce byte-for-byte from
this build. Historical synthetic-price exports move to archive/, and the
README's hand-copied Planner revision is dropped in favour of reading the pin
from Cargo.toml, since the copied one had gone stale.

The sort-key contract test now compiles the query instead of reading a
committed export, so it no longer depends on plans from an older Planner.

A certified companion fixture for the heap candidates is NOT included: the
issue-754 generator cannot supply one. Scoped evidence needs
topk_selected_lower_bound > topk_excluded_upper_bound, and under the
generator's own per-series domain the third- and fourth-ranked series overlap
in both groups, while these queries evaluate in real_time scope so the bound
must hold across the whole validity window. That needs a generator with
non-overlapping per-series domains.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@zzylol zzylol closed this Sep 28, 2026
@zzylol
zzylol deleted the test/promql-exact-function-coverage branch September 28, 2026 17:23
@zzylol
zzylol restored the test/promql-exact-function-coverage branch September 28, 2026 17:23
@zzylol zzylol reopened this Sep 28, 2026
@zzylol zzylol changed the title test: validate individual and ensemble physical candidate DAGs before pricing test: assert Level-1 candidate membership and account for every rejection Sep 28, 2026
zzylol added a commit that referenced this pull request Sep 28, 2026
An earlier revision claimed every recorded defect was a regression introduced
inside the #737 -> #728 stack, on the grounds that main is green on these
paths. Diffing residual_nodes against main does not support that: the function
is substantively identical there, and the stack changed only how the accuracy
target is derived and the wording of the error message.

main is green because its tests never plan these queries, not because the code
is correct. That is the more worrying reading, so it should be the recorded one.

The fragment mismatch is fixed on main by #781: a context-derived leaf schema
was compared for equality against an isolated re-parse that cannot reproduce
it, and open schemas are now compared by containment.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Sep 28, 2026
The workload-snapshot version was restated at every site that produced or
consumed a snapshot: the backend's check, two o11y tools, a shared-workload
tool, the shipped example snapshots and their tests. Nothing tied them
together, so they drifted. On the #728 stack two tools still overwrote the
version with a stale literal of 2 while the backend had moved to 3, which meant
a snapshot built from a current template was relabelled to an old version and
then rejected by the very backend that had produced the template.

Name it once, as WORKLOAD_SNAPSHOT_VERSION beside the field it governs, and use
it for both the check and its message. Producers are written in Rust, Python
and JSON and cannot share a constant, so tie them together with a test instead:
the shipped snapshots must declare that version, and each tool must reject
anything else and must not relabel what it is handed. A tool changes a
snapshot's content, not its schema, so it has no business restating the
version; that is precisely how the literal went stale.

Bumping the schema is now one edit plus whatever the test reports, rather than
a literal that some producers follow and others quietly do not.

Verified by mutation: raising the constant to 4 fails both tests, and
reintroducing a `snapshot["snapshot_version"] =` write into a tool fails with
the message naming that tool.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…tion

Enumerate the supported candidate inventory before pricing for ten
individual queries and three ensembles. Check typed DAG dependencies,
windows, grouping, Rate sort expressions, materialization boundaries,
dataset-bound SDS definitions and shared producer identity. The strict
fixture rejects heap candidates without certified accuracy evidence; the
certified companion proves those families bind. Nothing is priced.

Rebuilt from the #728 head (6042bb0) onto the ERP-free #761.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Sep 29, 2026
Adds #728's Level-1 tests, which the split base now contains.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@zzylol
zzylol force-pushed the test/promql-exact-function-coverage branch from 6042bb0 to 649fe08 Compare September 29, 2026 01:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant